Skip to content

fix(pack): keep blobs larger than 512 MiB readable - #163

Merged
wesm merged 2 commits into
kenn-io:mainfrom
hansn74:fix/pack-oversized-single-segment-frames
Oct 8, 2026
Merged

wesm merged 2 commits into
kenn-io:mainfrom
hansn74:fix/pack-oversized-single-segment-frames

Conversation

@hansn74

@hansn74 hansn74 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Existing blobs larger than 512 MiB become readable again. The previous writer emitted single-segment zstd frames whose effective window equals the full blob size, exceeding both the buffered decoder and default streaming reader's 512 MiB window limits.

On 64-bit systems, both readers now allow windows up to the format's 4 GiB raw-length ceiling. The 512 MiB window ceiling remains on 32-bit systems, where larger windows can overflow the codec's integer buffer calculations. This raises the default streaming memory allowance: reading an existing single-segment frame can require a decoder window as large as the blob. Callers can set a lower streaming window limit explicitly; packstore's bounded read APIs retain their policy limits, while Store.Open uses the raised default.

New blobs over 512 MiB use an explicit window within the klauspost decoder's default limit. Smaller blobs retain their existing framing, and existing packs need no rewrite.

encodeFrame compresses every blob with WithSingleSegment(true). A
single-segment frame carries no window descriptor, so a decoder derives its
window from the frame content size, while normalizeReaderLimits and the shared
zstdDec both cap windows at zstd.MaxWindowSize (512 MiB). Blobs above that are
therefore written in a form this package refuses to read:

    pack: corrupt: zstd decode: window size exceeded

The data is intact; it decodes correctly once a larger window is allowed.

Reading is what unblocks an existing repository: backup walks the hash-map
chain through such a blob before writing, so once one exists no further
snapshots can be created at all.

- reader.go: default WindowBytes ceiling raised from zstd.MaxWindowSize to
  MaxRawLen. Decoder memory stays bounded by WithDecoderMaxMemory in openBlob.
- frame.go: the shared whole-blob decoder gets the same allowance.
- frame.go: blobs over zstd.MaxWindowSize are encoded without single-segment
  framing, so they carry an explicit window descriptor and stay within what a
  stock decoder accepts.

Reaching the limit needs an unusually large blob. The case that surfaced this
was 681.7 MiB, produced when VACUUM rewrote a SQLite archive into one
contiguous run; neighbouring blobs in the same pack were 1.6 and 2.2 MiB.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@roborev-ci

roborev-ci Bot commented Oct 6, 2026

Copy link
Copy Markdown

roborev: Combined Review (60c9ba3)

Verdict: No findings at or above medium severity.


Reviewers: 2x codex, codex (security) | Synthesis: codex, 6s | Total: 1m46s

@wesm wesm self-assigned this Oct 6, 2026
@wesm

wesm commented Oct 6, 2026

Copy link
Copy Markdown
Member

Thanks for catching, reviewing now

@mariusvniekerk mariusvniekerk self-assigned this Oct 7, 2026
@wesm

wesm commented Oct 7, 2026

Copy link
Copy Markdown
Member

i have fixes prpepared

Existing repositories need coverage for frames written before the encoder
fix. A round trip through the new encoder cannot detect regressions in
those reads, and four huge payloads imposed unnecessary memory cost on CI.

Allowing 4 GiB windows also lets valid legacy frames overflow the codec's
integer buffer calculations on 32-bit systems. Preserve the previous
512 MiB ceiling there. On 64-bit systems, reading existing packs requires
the larger allowance and its corresponding memory cost.

Generated with Codex
Co-authored-by: Codex <noreply@openai.com>
@roborev-ci

roborev-ci Bot commented Oct 7, 2026

Copy link
Copy Markdown

roborev: Combined Review (de5032b)

Verdict: No findings at or above medium severity.


Reviewers: 2x codex, codex (security) | Synthesis: codex, 6s | Total: 2m20s

@wesm
wesm merged commit 2f2ed4c into kenn-io:main Oct 8, 2026
9 checks passed
wesm added a commit to kenn-io/msgvault that referenced this pull request Oct 8, 2026
Update kit to v0.32.2 so backup and restore can read existing pack blobs larger than 512 MiB on 64-bit systems. Newly written oversized blobs use a smaller zstd window, and existing packs need no rewrite.

Reading an old single-segment frame can require memory proportional to its full blob size. Explicit streaming limits remain enforced, and 32-bit decoders retain their 512 MiB window ceiling.

Includes the fix from [kit #163](kenn-io/kit#163).


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
wesm added a commit to kenn-io/docbank that referenced this pull request Oct 8, 2026
Update kit to v0.32.2 so backup and restore can read existing pack blobs larger than 512 MiB on 64-bit systems. Newly written oversized blobs use a smaller zstd window, and existing packs need no rewrite.

Reading an old single-segment frame can require memory proportional to its full blob size. Explicit streaming limits remain enforced, and 32-bit decoders retain their 512 MiB window ceiling.

Includes the fix from [kit #163](kenn-io/kit#163).


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
mariusvniekerk added a commit that referenced this pull request Oct 8, 2026
The tests from #163 for blobs larger than 512 MiB built real payloads of that size. One needed about 1.6 GB of memory in every ordinary CI run, on all three operating systems. To limit that cost, the file was excluded from race builds, so the race detector never ran these tests. This change improves the developer experience: it replaces those tests with ones that use small data, need little memory, and also run under the race detector.

- **Encoder:** the test lowers the single-segment cutoff to 64 KiB. The cutoff is now a package variable, `maxSingleSegmentLen`. A blob at the cutoff stays single-segment. A blob one byte over gets an explicit window.
- **Reader:** the test takes a 1 MiB single-segment frame and changes its header to claim 512 MiB + 1. `decodeFrame` and `OpenBlob` both check the window against that header before they decode. So the test shows that they accept the frame, without allocating 512 MiB. It skips on 32-bit, where `TestReaderRejectsOversized32BitWindow` already covers the 512 MiB cap.

Tradeoff: no test now decodes a real blob over 512 MiB from end to end. I reverted each of the three parts of the #163 fix in turn: the single-segment cutoff, the buffered decoder's window, and the default streaming window. Each revert makes the matching test fail.

<sup>generated by a clanker</sup>


Co-authored-by: Marius van Niekerk <mariusvniekerk@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants